Papers by Saif M. Mohammad
BRIGHTER: BRIdging the Gap in Human-Annotated Textual Emotion Recognition Datasets for 28 Languages (2025.acl-long)
Copied to clipboard
Shamsuddeen Hassan Muhammad, Nedjma Ousidhoum, Idris Abdulmumin, Jan Philip Wahle, Terry Ruas, Meriem Beloucif, Christine de Kock, Nirmal Surange, Daniela Teodorescu, Ibrahim Said Ahmad, David Ifeoluwa Adelani, Alham Fikri Aji, Felermino D. M. A. Ali, Ilseyar Alimova, Vladimir Araujo, Nikolay Babakov, Naomi Baes, Ana-Maria Bucur, Andiswa Bukula, Guanqun Cao, Rodrigo Tufiño, Rendi Chevi, Chiamaka Ijeoma Chukwuneke, Alexandra Ciobotaru, Daryna Dementieva, Murja Sani Gadanya, Robert Geislinger, Bela Gipp, Oumaima Hourrane, Oana Ignat, Falalu Ibrahim Lawan, Rooweither Mabuya, Rahmad Mahendra, Vukosi Marivate, Alexander Panchenko, Andrew Piper, Charles Henrique Porto Ferreira, Vitaly Protasov, Samuel Rutunda, Manish Shrivastava, Aura Cristina Udrea, Lilian Diana Awuor Wanzare, Sophie Wu, Florian Valentin Wunderlich, Hanif Muhammad Zhafran, Tianhui Zhang, Yi Zhou, Saif M. Mohammad
| Challenge: | Emotion recognition is an umbrella term for several NLP tasks, but most work on high-resource languages has focused on low-resourced languages. |
| Approach: | They propose to use emotion recognition to describe perceived emotions in 28 different languages and across several domains to identify and annotate the datasets. |
| Outcome: | The proposed datasets cover low-resource languages from Africa, Asia, Eastern Europe, and Latin America, with instances labeled by fluent speakers. |
Building Better: Avoiding Pitfalls in Developing Language Resources when Data is Scarce (2025.acl-long)
Copied to clipboard
| Challenge: | Language is a powerful means of communication and should be regarded as more than just a collection of tokens. |
| Approach: | They collect feedback from individuals directly involved in and impacted by NLP artefacts for medium- and low-resource languages and highlight key issues related to data quality, cultural appropriateness and ethics of common annotation practices. |
| Outcome: | The findings highlight key issues related to data quality, cultural appropriateness, and ethics of common annotation practices. |
Geographic Citation Gaps in NLP Research (2022.emnlp-main)
Copied to clipboard
| Challenge: | a vast number of papers accepted at top NLP venues come from a handful of western countries and (lately) China. |
| Approach: | They ask researchers to examine the relationship between geographical location and publication success . they use a dataset of 70,000 papers from the ACL Anthology to examine their citation network . |
| Outcome: | The proposed dataset of 70,000 papers from the ACL Anthology shows that there are substantial geographical disparities in paper acceptance and citations . |
PoKi: A Large Dataset of Poems by Children (2020.lrec-1)
Copied to clipboard
| Challenge: | a new corpus of child-written texts is available for study of child language . authors use non-parametric regressions to model developmental differences from early childhood to late-adolescence . |
| Approach: | They propose to analyze 62 thousand child-written poems written by children from grades 1 to 12 . they use non-parametric regressions to model developmental differences from early childhood to late-adolescence . |
| Outcome: | The proposed corpus includes about 62 thousand poems written by children from grades 1 to 12 . results show decreases in valence that are especially pronounced during mid-adolescence . |
Citation Amnesia: On The Recency Bias of NLP and Other Academic Fields (2025.coling-main)
Copied to clipboard
| Challenge: | citation age is a key factor in determining whether older works are cited in scientific journals or not. |
| Approach: | They examine the tendency of NLP to cite older work across 20 fields of study over 43 years (1980–2023) . they put NLP’s propensity to citation older work in the context of these 20 other fields to see whether differences can be observed . |
| Outcome: | The trend is strongest in NLP and ML research (-12.8% and -5.5% in citation age from previous peaks) |
Gender Gap in Natural Language Processing Research: Disparities in Authorship and Citations (2020.acl-main)
Copied to clipboard
| Challenge: | Disparities in authorship and citations across gender can have adverse consequences . Historically, gender has been considered binary (male and female), immutable (cannot change), and physiological (mapped to biological sex). |
| Approach: | They examine female first author percentages and citations to papers in natural language processing . they find that only about 29% of first authors are female and only about 25% of last authors are male . |
| Outcome: | The authors show that only about 29% of first authors are female and only about 25% of last authors are male . the authors argue that gender gaps are unfair and need to be addressed . |
Examining Citations of Natural Language Processing Literature (2020.acl-main)
Copied to clipboard
| Challenge: | citations of NLP papers have decreased in recent years, but long papers get three times as many citation as short papers . citation data from the ACL Anthology and Google Scholar can be used to understand the field and quantify the impact of different types of papers. |
| Approach: | They extract data from the ACL Anthology and Google Scholar to examine trends in citations of NLP papers. |
| Outcome: | The results show that only about 56% of the papers in AA are cited ten or more times . CL Journal has the most cited papers, but its citation dominance has lessened . |
NLP Scholar: A Dataset for Examining the State of NLP Research (2020.lrec-1)
Copied to clipboard
| Challenge: | Google Scholar is the largest web search engine for academic literature and provides access to rich metadata associated with the papers. |
| Approach: | They extracted citation information from the ACL Anthology (AA) for about 44 thousand NLP papers and identified authors who published at least three papers there. |
| Outcome: | The ACL Anthology (AA) is the largest repository of articles on Natural Language Processing (NLP). |
WordWars: A Dataset to Examine the Natural Selection of Words (2020.lrec-1)
Copied to clipboard
| Challenge: | a growing body of work on how word meaning changes over time is mutation . a new dataset, WordWars, explores how word success changes over the time . |
| Approach: | They analyze a dataset of 5000 English words in synsets and examine natural selection . they find frequency, length, and concreteness all impact natural selection, they say . |
| Outcome: | a new dataset shows that one third of the synsets undergo a change in the predominant word in this time period. |
DimABSA: Building Multilingual and Multidomain Datasets for Dimensional Aspect-Based Sentiment Analysis (2026.acl-long)
Copied to clipboard
Lung-Hao Lee, Liang-Chih Yu, Natalia V Loukachevitch, Ilseyar Alimova, Alexander Panchenko, Tzu-Mi Lin, Zhe-Yu Xu, Jian-Yu Zhou, Guangmin Zheng, Jin Wang, Sharanya Awasthi, Jonas Becker, Jan Philip Wahle, Terry Ruas, Shamsuddeen Hassan Muhammad, Saif M. Mohammad
| Challenge: | Existing ABSA research relies on coarse-grained categorical labels, which limits its ability to capture nuanced affective states. |
| Approach: | They propose a dimensional approach that represents sentiment with continuous valence–arousal (VA) scores, enabling fine-grained analysis at both the aspect and sentiment levels. |
| Outcome: | The proposed approach represents sentiment with continuous valence–arousal (VA) scores, enabling fine-grained analysis at both the aspect and sentiment levels. |
What Media Frames Reveal About Stance: A Dataset and Study about Memes in Climate Change Discourse (2025.findings-emnlp)
Copied to clipboard
| Challenge: | Media framing is a method of shaping public perceptions of issues, but the interaction between stance and media frame remains unexplored. |
| Approach: | They propose to use a dataset of climate-change memes annotated with stance and media frames to conceptualize and computationally explore this interaction. |
| Outcome: | The proposed dataset includes 1,184 climate-change memes sourced from 47 subreddits and enables analysis of frame prominence over time and communities. |
Words of Warmth: Trust and Sociability Norms for over 26k English Words (2025.acl-long)
Copied to clipboard
| Challenge: | Social psychologists have shown that Warmth (W) and Competence (C) are the primary dimensions along which we assess other people and groups. |
| Approach: | They propose a repository of word–warmth and word–trust associations for over 26k English words. |
| Outcome: | The proposed lexicon enables bias and stereotype research through case studies on target entities. |
Annotating Dimensions of Social Perception in Text: A Sentence-Level Dataset of Warmth and Competence (2026.acl-long)
Copied to clipboard
| Challenge: | *Warmth* (W) and *Competence (C) are central dimensions along which people evaluate individuals and social groups. |
| Approach: | They propose a first sentence-level dataset annotated for warmth and competence . they analyze sentences that express attitudes and opinions about individuals or social groups . |
| Outcome: | The first sentence-level dataset annotated for warmth and competence is presented in this paper. |
The Nature of NLP: Analyzing Contributions in NLP Papers (2025.acl-long)
Copied to clipboard
| Challenge: | despite this, what constitutes NLP research remains debated . |
| Approach: | They propose a taxonomy of research contributions and introduce a task of automatically identifying contribution statements and classifying their types from NLP research papers. |
| Outcome: | The proposed model analyzes 29k NLP research papers to understand their contributions . |
Ruddit: Norms of Offensiveness for English Reddit Comments (2021.acl-long)
Copied to clipboard
| Challenge: | Existing methods to detect offensive language have been limited by categorical labels . however, there are several challenges in the detection of such content . |
| Approach: | They analyze Reddit comments with fine-grained, real-valued offensiveness scores . they evaluate the ability of widely-used neural models to predict offensiveness . |
| Outcome: | The proposed method produces highly reliable offensiveness scores and can predict scores on reddit comments. |
NLP Scholar: An Interactive Visual Explorer for Natural Language Processing Literature (2020.acl-demos)
Copied to clipboard
| Challenge: | aCL Anthology and Google Scholar provide a single dataset of NLP papers and their meta-information . authors describe interactive visualizations that present various aspects of the data . |
| Approach: | They propose to use citation data from the ACL Anthology and Google Scholar to create a unified dataset of NLP papers and their meta-information. |
| Outcome: | The proposed dataset includes papers published in the area of their interest and by specified authors. |
SOLO: A Corpus of Tweets for Examining the State of Being Alone (2020.lrec-1)
Copied to clipboard
| Challenge: | Psychologists distinguish between the concept of solitude, a positive state of voluntary aloneness, and the concept 'loneliness', characterized as dissatisfaction with the quality of one’s social interactions. |
| Approach: | They present a corpus of over 4 million tweets with query terms solitude, lonely, and loneliness. |
| Outcome: | The proposed analysis analyzes over 4 million tweets with the terms solitude, lonely, and loneliness. |
The Language of Interoception: Examining Embodiment and Emotion Through a Corpus of Body Part Mentions (2025.findings-emnlp)
Copied to clipboard
| Challenge: | 5% to 10% of posts include body part mentions in English text . text containing BPMs tends to be more emotionally charged, even when the BPM is not used to describe a physical reaction to the emotion in the text. |
| Approach: | They create corpora of body part mentions in online English text with human annotations for the emotions of the person whose body part is mentioned. |
| Outcome: | The proposed study is the first to investigate the connection between emotion, embodiment, and everyday language in a large sample of natural language data. |
Tweet Emotion Dynamics: Emotion Word Usage in Tweets from US and Canada (2022.lrec-1)
Copied to clipboard
| Challenge: | a dataset of 45 million geo-located tweets from the US and Canada is used to analyze emotions . early work identified tweets as a crucial indicator of public sentiment . |
| Approach: | They propose a dataset of more than 45 million geo-located tweets from US and Canada . they also introduce Tweet Emotion Dynamics (TED) metrics to capture patterns of emotions associated with tweets . |
| Outcome: | The proposed dataset includes more than 45 million geo-located tweets from US and Canada . it shows that Canadian tweets tend to have higher valence, lower arousal, and higher dominance than the US tweets . |